Distributed Application Checkpointing for Replicated State Machines

نویسندگان

چکیده

Application checkpointing is a widely used recovery mechanism that consists of saving an application's state periodically to be in case failure. In this study we investigate the utilisation distributed for replicated machines. Conventionally, machines, information stored way each replicas or separately single instance. Applying provides means adjust level fault tolerance approach by giving away from time. We use local cluster and cloud environment examine effects simple machine example compare results with conventional approaches. As expected, gains memory consumption utilise different levels while performing worse terms

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

CASPaxos: Replicated State Machines without logs

CASPaxos is a replicated state machine (RSM) protocol, an extension of Synod. Unlike Raft and Multi-Paxos, it doesn’t use leader election and log replication, thus avoiding associated complexity. Its symmetric peer-to-peer approach achieves optimal commit latency in wide-area networks and doesn’t cause transient unavailability when any bN−1 2 c of N nodes crash. The lightweight nature of CASPax...

متن کامل

Fast Replicated State Machines Over Partitionable Networks

This paper presents an implementationof replicated state machines in asynchronous distributed environments prone to node failures and network partitions. This implementation has several appealing properties: It guarantees that progress will be made whenever a majority of replicas can communicate with each other; it allows minority partitions to continue providing service for idempotent requests...

متن کامل

Mencius: Building Efficient Replicated State Machines for WANs

We present a protocol for general state machine replication – a method that provides strong consistency – that has high performance in a wide-area network. In particular, our protocol Mencius has high throughput under high client load and low latency under low client load even under changing wide-area network environment and client load. We develop our protocol as a derivation from the well-kno...

متن کامل

Application controlled checkpointing coordination for fault-tolerant distributed computing systems

In order to provide fault tolerance for distributed systems, the checkpointing technique has widely been used and many researches have been performed to reduce the overhead of check-pointing coordination. In this paper, we present a new checkpointing coordination scheme in which the application controls the coordination activity by utilizing the communication pattern of the application program....

متن کامل

Canonical finite state machines for distributed systems

There has been much interest in testing from finite state machines (FSMs) as a result of their suitability for modelling or specifying state-based systems. Where there are multiple ports/interfaces a multi-port FSM is used and in testing a tester is placed at each port. If the testers cannot communicate with one another directly and there is no global clock then we are testing in the distribute...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

ژورنال

عنوان ژورنال: Scalable Computing: Practice and Experience

سال: 2021

ISSN: ['1895-1767']

DOI: https://doi.org/10.12694/scpe.v22i1.1840